Papers with vision-language models pre-

1 papers
Vision-Language Pre-Training for Multimodal Aspect-Based Sentiment Analysis (2022.acl-long)

Copied to clipboard

Challenge: Existing approaches to multimodal Aspect-Based Sentiment Analysis (MABSA) ignore crossmodalalignment and use pre-trained visual and textual models.
Approach: They propose a multimodal multimodal encoder-decoder framework for MABSA that uses a unified multimodal decoder architecture for all the pretrainingand downstream tasks.
Outcome: The proposed framework outperforms state-of-the-art approaches on three MABSA subtasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations